classroom environment
Speech Separation for Hearing-Impaired Children in the Classroom
Olalere, Feyisayo, van der Heijden, Kiki, Stronks, H. Christiaan, Briaire, Jeroen, Frijns, Johan H. M., Güçlütürk, Yagmur
The process includes simulating room and listener acoustic properties (A), modeling talkers' movement trajectories (B), and synthesizing classroom speech mixtures (C). The numbers (1) - (5) correspond to the steps itemized in section II-B more challenging and reflective of classroom acoustics. The separation model is trained to output time-domain waveforms for each speaker with no interference from the other speaker or background noise. This setup enables the model to not only separate overlapping speech, but also to preserve spatial distinctions associated with each moving source. B. Simulation of Overlapping Speech for Classroom Conditions To capture the reverberant and spatial characteristics typical of classroom environments, we developed a spatialization pipeline for generating training and evaluation data (see Fig.1). This pipeline consists of five main components, which are explained below in detail: 1) Simulation of room impulse responses (RIRs) 2) Application of head-related impulse responses (HRIRs) 3) Generation of binaural room impulse responses (BRIRs) 4) Modeling of talkers' movement trajectories 5) Synthesis of the classroom speech data 1) Room Impulse Responses: To simulate naturalistic reverberant classroom acoustics, we generated RIRs that capture direct sound, early reflections, and reverberation or echo. These RIRs were used to spatialize source signals in simulated classroom environments with varying geometry, reverberation, and source-listener distances. We used the Pyroomacoustics Python package [35], which implements the image source method to model sound propagation in rectangular (shoebox) rooms. A total of 30 classrooms were simulated, with dimensions randomly sampled from a range of 8.5 8.5 3 m to 10 10 3.5 m (length width height), reflecting typical U.S. classroom sizes [36], [37].
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
Attia, Ahmed Adel, Liu, Jing, Espy-Wilson, Carl
The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Public classroom datasets remain limited, and the lack of a dedicated classroom noise corpus prevents the use of standard data augmentation techniques. In this paper, we introduce a scalable methodology for synthesizing classroom noise using game engines, a framework that extends to other domains. Using this methodology, we present SimClass, a dataset that includes both a synthesized classroom noise corpus and a simulated classroom speech dataset. The speech data is generated by pairing a public children's speech corpus with Y ouTube lecture videos to approximate real classroom interactions in clean conditions. Our experiments on clean and noisy speech demonstrate that SimClass closely approximates real classroom speech, making it a valuable resource for developing robust speech recognition and enhancement models.
Multi-Stage Speaker Diarization for Noisy Classrooms
Khan, Ali Sartaz, Ogunremi, Tolulope, Attia, Ahmed Adel, Demszky, Dorottya
Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording quality, high levels of background noise, overlapping speech, and the difficulty of accurately capturing children's voices. This study investigates the effectiveness of multi-stage diarization models using Nvidia's NeMo diarization pipeline. We assess the impact of denoising on diarization accuracy and compare various voice activity detection (VAD) models, including self-supervised transformer-based frame-wise VAD models. We also explore a hybrid VAD approach that integrates Automatic Speech Recognition (ASR) word-level timestamps with frame-level VAD predictions. We conduct experiments using two datasets from English speaking classrooms to separate teacher vs. student speech and to separate all speakers. Our results show that denoising significantly improves the Diarization Error Rate (DER) by reducing the rate of missed speech. Additionally, training on both denoised and noisy datasets leads to substantial performance gains in noisy conditions. The hybrid VAD model leads to further improvements in speech detection, achieving a DER as low as 17% in teacher-student experiments and 45% in all-speaker experiments. However, we also identified trade-offs between voice activity detection and speaker confusion. Overall, our study highlights the effectiveness of multi-stage diarization models and integrating ASR-based information for enhancing speaker diarization in noisy classroom environments.
Examining the Role of LLM-Driven Interactions on Attention and Cognitive Engagement in Virtual Classrooms
Ozdel, Suleyman, Sarpkaya, Can, Bozkir, Efe, Gao, Hong, Kasneci, Enkelejda
Transforming educational technologies through the integration of large language models (LLMs) and virtual reality (VR) offers the potential for immersive and interactive learning experiences. However, the effects of LLMs on user engagement and attention in educational environments remain open questions. In this study, we utilized a fully LLM-driven virtual learning environment, where peers and teachers were LLM-driven, to examine how students behaved in such settings. Specifically, we investigate how peer question-asking behaviors influenced student engagement, attention, cognitive load, and learning outcomes and found that, in conditions where LLM-driven peer learners asked questions, students exhibited more targeted visual scanpaths, with their attention directed toward the learning content, particularly in complex subjects. Our results suggest that peer questions did not introduce extraneous cognitive load directly, as the cognitive load is strongly correlated with increased attention to the learning material. Considering these findings, we provide design recommendations for optimizing VR learning spaces.
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
Attia, Ahmed Adel, Demszky, Dorottya, Ogunremi, Tolulope, Liu, Jing, Espy-Wilson, Carol
Creating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We show that CPT is a powerful tool in that regard and reduces the Word Error Rate (WER) of Wav2vec2.0-based models by upwards of 10%. More specifically, CPT improves the model's robustness to different noises, microphones and classroom conditions.
Mandarin Language Learners Get A Boost From AI
IBM Research and Rensselaer Polytechnic Institute (RPI) are collaborating on a new approach to help students learn Mandarin. The strategy pairs an AI-powered assistant with an immersive classroom environment that has not been used previously for language instruction. The classroom, called the Cognitive Immersive Room (CIR), makes students feel as though they are in restaurant in China, a garden, or a Tai Chi class, where they can practice speaking Mandarin with an AI chat agent. The CIR was developed by the Cognitive and Immersive Systems Lab (CISL), a research collaboration between IBM Research and RPI. When learning a new language, especially one as difficult as Mandarin, it's important that students have many opportunities to speak and practice their conversational skills.
Transforming HE through machine learning
Undoubtedly, the digital revolution has transformed nearly every industry. At the forefront of this transformation are Artificial Intelligence (AI) and Machine Learning (ML). While many industries, like transportation and retail, have become leaders in adopting emerging technology to improve their business models, the higher education industry has fallen behind. This gap is present in some lower levels of schooling as well, but, historically, we've seen that larger university settings have encountered more challenges in the path to adoption. Prior to entering higher education, students are exposed to, and leverage, various forms of technology.
UK education expert dismisses 'Minecraft' as a 'gimmick'
After offering teachers early access to Minecraft: Education Edition this summer, Microsoft's classroom-friendly version of the immensely popular sandbox game was formally launched at the beginning of November. Not everyone is keen on Minecraft being used as a teaching tool, though, and ahead of Microsoft's UK launch event tomorrow, behavior expert for the government's Department for Education Tom Bennett has voiced his skepticism to The Times. "I am not a fan of Minecraft in lessons. This smacks to me of another gimmick which will get in the way of children actually learning," Bennett said. "Removing these gimmicky aspects of education is one of the biggest tasks facing us as teachers. We need to drain the swamp of gimmicks," he continued, mimicking some recent rhetoric from US President-elect Trump.